Papers with reward estimation
Out of Distribution, Out of Luck: Process Rewards Misguide Reasoning Models (2026.eacl-short)
Copied to clipboard
| Challenge: | 80% of reasoning model outputs respond to formatting artifacts rather than mathematical content. |
| Approach: | They evaluate process reward models that provide step-level feedback during inference . they identify distinct reward prediction patterns that differentiate reasoning from non-reasoning model outputs . |
| Outcome: | The proposed model fails to enhance and sometimes degrade reasoning model performance. |
PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing approaches to align large language models with human preferences are limited by their large-scale annotation and prone to reward overoptimization. |
| Approach: | They propose a training paradigm that integrates three complementary strategies to address these challenges by reformulating question–answer pairs into preference-task instructions, averaging the rewards aggregated from diverse preference- task instructions for each sample, and a balancing outputs from the value head under different dropout rates. |
| Outcome: | Experiments on public datasets show that PIRA improves performance considerably, enhances generalization, and effectively mitigates reward overoptimization. |
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are prone to hallucination, especially during multihop tasks. |
| Approach: | They propose a hierarchical, erroraware discriminative PRM that classifies math errors at each step and combines finegrained signals to estimate step correctness. |
| Outcome: | The proposed model outperforms the prior best in a new stateof-theart PRMScore of 67.7 on a 400Ksample dataset . |